Papers with Automated Essay Scoring
Beyond the Gold Standard in Analytic Automated Essay Scoring (2025.acl-srw)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) is a new approach to assessing writing practice . traditional holistic scoring methods are not reliable and lack formative feedback in the classroom. |
| Approach: | They propose to combine analytic and holistic AES to create a system that learns from individual raters instead of gold standard labels. |
| Outcome: | The proposed system learns from individual raters instead of gold standard labels. |
Multi-task Learning for Automated Essay Scoring with Sentiment Analysis (2020.aacl-srw)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) is a process that aims to alleviate the workload of graders and improve the feedback cycle in educational systems. |
| Approach: | They propose to combine two tasks, sentiment analysis and AES by utilizing multi-task learning to combine sentiment features extracted from opinion expressions. |
| Outcome: | The proposed model produces a QWK of 0.763 on the Automated StudentAssessment Prize (ASAP) benchmark. |
Neural Automated Essay Scoring and Coherence Modeling for Adversarially Crafted Input (N18-1)
Copied to clipboard
| Challenge: | Existing approaches to Automated Essay Scoring (AES) are not well-suited to capture adversarially crafted input of grammatical but incoherent sequences of sentences. |
| Approach: | They propose a neural model of local coherence that can effectively learn connectedness features between sentences. |
| Outcome: | The proposed approach strengthens the validity of neural essay scoring models. |
Qayyem: A Real-time Platform for Scoring Proficiency of Arabic Essays (2026.acl-demo)
Copied to clipboard
| Challenge: | Existing Arabic writing technologies primarily use a single quality score for essays, but there is limited support for Arabic AES. |
| Approach: | They propose a Web-based platform that integrates Arabic AES workflows with a user-friendly interface. |
| Outcome: | The proposed system integrates with existing Arabic scoring systems and provides a user-friendly interface. |
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments (2025.acl-industry)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) systems aim to evaluate the quality of candidate writing using computational methods. |
| Approach: | They propose a model that assigns a confidence score to each automated score to ensure it meets high reliability standards. |
| Outcome: | The proposed model achieves an F1 score of 0.97 and releases 47% of predicted scores with 100% CEFR agreement and 99% with at least 95% CEFR agreeance compared to the standalone model where all predicted scores are released. |
Enhancing Automated Essay Scoring Performance via Fine-tuning Pre-trained Language Models with Combination of Regression and Ranking (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work on sentence prediction tasks uses shallow neural networks to learn essay representations and constrain calculated scores with regression loss or ranking loss. |
| Approach: | They propose to use a pre-trained language model to learn text representations first and then to constrain the scores with regression loss or ranking loss. |
| Outcome: | The proposed model outperforms state-of-the-art models on the Automated Student Assessment Prize dataset. |
LAILA: A Large Trait-Based Dataset for Arabic Automated Essay Scoring (2026.eacl-long)
Copied to clipboard
May Bashendy, Walid Massoud, Sohaila Eltanbouly, Salam Albatarni, Marwan Sayed, Abrar Abir, Houda Bouamor, Tamer Elsayed
| Challenge: | Existing Arabic resources are small in scale and lack trait-specific annotations. |
| Approach: | They propose to use LAILA to build a large Arabic AES dataset with holistic and trait-specific annotations of seven writing proficiency traits. |
| Outcome: | The LAILA dataset comprises 7,859 essays annotated with holistic and trait-specific scores on seven dimensions: relevance, organization, vocabulary, style, development, mechanics, and grammar. |
TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay Scoring (2023.emnlp-main)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) aims to automatically assess the quality of essays. |
| Approach: | They propose to use a corpus of 6.5k essays collected in the context of the Test de Connaissance du Français (TCF) certification exam to foster the development of AES for French. |
| Outcome: | The proposed system can assess the quality of essays in a language certification exam using a corpus of 6.5k essays collected in the TCFLE-8 exam. |
On the Use of Bert for Automated Essay Scoring: Joint Learning of Multi-Scale Essay Representation (2022.naacl-main)
Copied to clipboard
| Challenge: | Pre-trained models have not been used to outperform other deep learning models such as CNN in Automated Essay Scoring (AES). |
| Approach: | They propose a novel multi-scale essay representation for BERT that can be jointly learned . they employ multiple losses and transfer learning from out-of-domain essays to further improve performance . |
| Outcome: | The proposed model outperforms existing models in the area of automated essay scoring . the proposed model generalizes well to the CommonLit Readability Prize data set . |
Representation-to-Creativity (R2C): Automated Holistic Scoring Model for Essay Creativity (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies on Automated Essay Scoring (AES) are limited. |
| Approach: | They propose a new essay rubric specifically designed for assessing creativity in essays . they use a ground truth data set to construct a self-supervised learning model . |
| Outcome: | The proposed model improves the assessment of creativity in essays by 58% compared to the current models. |
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models (2025.findings-acl)
Copied to clipboard
Jiamin Su, Yibo Yan, Fangteng Fu, Zhang Han, Jingheng Ye, Xiang Liu, Jiahao Huo, Huiyu Zhou, Xuming Hu
| Challenge: | Automated Essay Scoring (AES) systems face three major challenges: reliance on handcrafted features that limit generalizability, difficulty in capturing fine-grained traits like coherence and argumentation, and inability to handle multimodal contexts. |
| Approach: | They propose a multimodal benchmark to evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering. |
| Outcome: | The proposed system can evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering. |
PsyScore: A Psychometrically-Aware Framework for Trait-Adaptive Essay Scoring and ZPD-Scaffolded Feedback (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to Automated Essay Scoring (AES) treat scoring and feedback as separate components, resulting in fragmentation. |
| Approach: | They propose a psychometrically-aware framework that integrates diagnostic assessment with instructional scaffolding through a shared latent ability representation. |
| Outcome: | The proposed framework integrates diagnostic assessment with instructional scaffolding through a shared latent ability representation. |
Computing with Subjectivity Lexicons (2020.lrec-1)
Copied to clipboard
Caio L. M. Jeronimo, Claudio E. C. Campelo, Leandro Balby Marinho, Allan Sales, Adriano Veloso, Roberta Viola
| Challenge: | a new set of lexicons for expressing subjectivity in text documents is presented . lexiconics are useful resources for identifying semantics relevant to sentiment, emotion, personality, language bias, mood, and attitude. |
| Approach: | They propose a set of lexicons for expressing subjectivity in Brazilian Portuguese text documents . they use word embedding techniques to capture semantically related words to the ones in the lexicos . |
| Outcome: | The proposed lexicons represent different subjectivity dimensions and are more compact in number of terms. |
Beyond Agreement: Diagnosing the Rationale Alignment of Automated Essay Scoring Methods based on Linguistically-informed Counterfactuals (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing Automated Essay Scoring (AES) methods focus on sentence-level features, whereas Large Language Models (LLMs) are sensitive to conventions & accuracy, language complexity, and organization. |
| Approach: | They propose to use large language models to aid in decision-making . they propose to analyze the reasoning of neural models by analyzing sentence-level features. |
| Outcome: | The proposed method improves understanding of neural approaches to Automated Essay Scoring (AES) and can also apply to other domains seeking transparency in model-driven decisions. |
Improving Domain Generalization for Prompt-Aware Essay Scoring via Disentangled Representation Learning (2023.acl-long)
Copied to clipboard
| Challenge: | Existing AES models are either prompt-specific or prompt-adaptive and cannot generalize well on “unseen” prompts. |
| Approach: | They propose a prompt-aware neural AES model to extract comprehensive representation for essay scoring, including both prompt-invariant and prompt-specific features. |
| Outcome: | The proposed model extracts comprehensive representation for essay scoring, including both prompt-invariant and prompt-specific features. |
Aggregating Multiple Heuristic Signals as Supervision for Unsupervised Automated Essay Scoring (2023.acl-long)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) aims to evaluate the quality score of input essays without human intervention. |
| Approach: | They propose an unsupervised approach to evaluate the quality of input essays . they use multiple heuristic quality signals as pseudo-groundtruths to train a neural AES model . |
| Outcome: | The proposed approach achieves state-of-the-art performance on eight prompts of ASPA dataset compared with previous unsupervised methods . |
Mixture of Ordered Scoring Experts for Cross-prompt Essay Trait Scoring (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to automate essay scoring overlook critical information, authors say . evaluators often limit their performance to unseen topics, resulting in incomplete assessment perspectives. |
| Approach: | They propose a framework that integrates information from prompts and essays into an AES framework. |
| Outcome: | The proposed framework achieves state-of-the-art in cross-prompt scoring and multi-trait scoring on the ASAP++ dataset. |
Towards Explainable Chinese Native Learner Essay Fluency Assessment: Dataset, Tasks, and Method (2024.findings-emnlp)
Copied to clipboard
Xinshu Shen, Hongyi Wu, Yadong Zhang, Man Lan, Xiaopeng Bai, Shaoguang Mao, Yuanbin Wu, Xinlin Zhuang, Li Cai
| Challenge: | Existing GEC datasets in Chinese fail to consider specific grammatical error types and overlook cross-sentence grammamatical errors. |
| Approach: | They propose to use Chinese essay fluency assessment to assess essay fluencies along with coarse and fine-grained errors and corrections to improve explainability. |
| Outcome: | The proposed dataset encapsulates essay fluency scores along with both coarse and fine-grained errors and corrections. |
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment (2025.emnlp-main)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) systems attain near–human agreement on some public benchmarks, but real-world adoption is limited. |
| Approach: | They propose a distribution-free wrapper that equips any classifier with set-valued outputs enjoying formal coverage guarantees. |
| Outcome: | The proposed model achieves coverage targets while keeping prediction sets compact. |
MAPLE: A Meta-learning Framework for Cross-Prompt Essay Scoring (2026.findings-acl)
Copied to clipboard
| Challenge: | Current approaches to automate essay scoring (AES) treat each writing task as a separate task, resulting in inconsistent performance. |
| Approach: | They propose a meta-learning framework that leverages prototypical networks to learn transferable representations across different writing prompts. |
| Outcome: | The proposed framework outperforms baseline models on ELLIPSE and ASAP (English) and LAILA (Arabic) on three diverse datasets. |
Transformer-based Joint Modelling for Automatic Essay Scoring and Off-Topic Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies show that automated essay scoring systems assign lower grades to irrelevant responses. |
| Approach: | They propose an unsupervised technique that jointly scores essays and detects off-topic essays. |
| Outcome: | The proposed method outperforms baseline and earlier conventional methods on two essay-scoring datasets in off-topic detection and on-topic scoring. |
Graph-Based Multi-Trait Essay Scoring (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on Automated Essay Scoring (AES) models essay as word sequence, but new approach uses graph-attention network approach to model essay traits. |
| Approach: | They propose a graph-attention network approach to automate essay scoring that models interactions among essay traits as a graphical graph. |
| Outcome: | The proposed approach outperforms competing approaches on the ASAP++ dataset . it allows for multiple-task scoring, allowing for more detailed feedback on essays . |